Genome Medicine
○ Springer Science and Business Media LLC
Preprints posted in the last 30 days, ranked by how well they match Genome Medicine's content profile, based on 183 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit.
Bermudez-Guzman, L.; Ramos-Esquivel, A.; Alpizar-Alpizar, W.
Show abstract
Early- and average-onset colorectal cancer (CRC) are separated at age 50, but whether this defines a biological threshold remains unclear. To clarify this, we identified molecular profiles in nine harmonised cBioPortal CRC cohorts (4,609 patients) by fitting Bernoulli mixture models to 31 repair-state, genomic-burden and gene-alteration features, excluding age, sex and tumour site, and compared their prevalence using <50/[≥]50 and decade-resolved groups. Four profiles captured conventional/CIN-like (P1), intermediate MSS (P2), KRAS/PI3K/APC-rich (P3) and hypermutated/MSI-high (P4) states along a left-to-right gradient. Although molecular identities remained stable, profile prevalence followed non-linear P1/P4 and opposing linear P2/P3 age trajectories. Profile-prevalence patterns did not track chronological proximity: profile composition at 30-39 differed from 50-59 but not clearly from 60-69. The age-50 threshold captured only 17.8% of decade-resolved deviance, whereas the optimal age-70 cut-off retained only 51.3%. Validation in 2,579 non-overlapping MSK-IMPACT patients (2,476 age-evaluable) reproduced molecular-feature patterns (r=0.97-0.98), age trajectories (r=0.92) and limited binary-threshold performance: age 50 and the optimal age-66 cut-off retained 12.7% and 45.1%, respectively. Thus, age reorganizes the prevalence of shared CRC states rather than defining a biological threshold at age 50.
Jiang, K.; Jarvis, J. N.
Show abstract
While progress has been made in identifying the true risk-driving single nucleotide polymorphisms (SNPS) on juvenile idiopathic arthritis (JIA) risk haplotypes, the affected cells and target genes largely remain unknown. We used data from a previously published massively parallel reporter assay (MPRA) to query human data in the Database of Immune Cell eQTLs (DICE) and the Gene-Tissue Expression (GTEx) database to identify affected cells and target genes of MPRA-identified SNPs in immune cells and relevant tissues. SNPs identified on MPRA were associated with gene expression levels in a broad range of immune cells in the DICE database, including CD4+ and CD8+ T lymphocytes, monocytes, NK cells, and B cells. MPRA-identified SNPs showed strong associations with gene expression in GTEx whole blood, spleen, and/or EBV-stimulated lymphocytes. Our data show the efficacy of combining MPRA and using human cells/tissue expression data to elucidate complex mechanisms driving genetic risk for JIA.
Multerer, K.; Atkinson, P.; Woods, L.; Tanigawa, Y.; Kellis, M.; Munkacsi, A.
Show abstract
Polygenic risk scores (PRS) assume additive SNP effects, yet genetic risk also arises from interactions between loci and environmental factors that contribute to broad-sense heritability. We developed an extended PRS (ePRS) framework for type 2 diabetes (T2D) that incorporates locus-by-locus non-additive effects beyond those captured by additive single-locus PRS or linkage disequilibrium (LD) tagging. These were modelled as cumulative burden (G+G; summed allele counts), statistical epistasis (GxG; allele count products), and gene-environment effects derived from cardiometabolic variables in electronic health records. Across 235,000 UK Biobank participants, five complementary ePRS models captured largely non-overlapping high-risk individuals, suggesting that a key to individual risk predictions comprise the inclusion of multiple interaction-driven biological components rather than a single signal. A composite score improved case detection beyond clinical predictors, including individuals within clinically normal ranges. These findings were generalized to celiac disease, with similar complementarity across models, with potential for clinical use pending prospective validation.
Farid, A. C.; Haldeman, S.; Otto, C.; DMello, A.; Tettelin, H.; Ratner, A. J.
Show abstract
Based on recent epidemiologic studies, Streptococcus agalactiae (Group B Streptococcus; GBS) sequence type (ST) 1010 is an emerging lineage now identified in multiple countries. We report the phylogenetic and genomic characteristics of a set of 55 GBS sequence type (ST) 1010 strains, as well as two newly described single-locus variants of ST1010. A core genome phylogeny suggests that ST1010 is closely related to both ST452 and the hypervirulent clonal complex (CC) 17 GBS lineage. Notably, we demonstrate that genes encoding two virulence determinants previously described as specific to CC17 GBS, the HvgA adhesin and the serine-rich repeat protein Srr2, are both present in ST1010 genomes. Srr2 is shared with members of ST452. High-level gentamicin resistance (HLGR) encoded on an IS256 mobile element, previously described in a small number of ST1010 isolates, is present in a distinct ST1010 subclade encompassing the majority of ST1010 isolates. The relationship between ST452 (serotype IV), ST1010 (serotype IV), and ST17 (serotype III) strains suggests that ST17 may have arisen from a serotype IV ancestor and later acquired the type III capsule locus. Taken together, these findings clarify the phylogenetic position of ST1010 and suggest sequential acquisition of virulence determinants and HLGR prior to its international emergence. IMPACT STATEMENTST1010 GBS has emerged internationally, with colonizing and invasive isolates described in the United States, Dominican Republic, Netherlands, and Italy. Using a core genome phylogeny and targeted detection of genomic regions, we demonstrate that ST1010 shares specific virulence determinants with the CC17 hypervirulent GBS lineage and that HLGR is confined to a specific numerically dominant subclade of ST1010. Our work spotlights the importance of future epidemiologic and genomic surveillance of ST1010 and related lineages. DATA SUMMARYPublicly available genomic data were used from three previously published studies (Laycock KM et al., McGee L et al., Khan UB et al.), as well as a set of newly sequenced GBS genomes from clinical strains originating in New York City (NYC). The corresponding accession numbers and detailed information for all strains are provided in the Table.
Viz-Lasheras, S.; Dacosta, A.; Rivero-Calle, I.; Martinon-Torres, F.; EUCLIDS, GENDRES, PERFORM, and DIAMONDS consortia, ; Gomez-Carballa, A.; Salas, A.
Show abstract
Accurate discrimination between viral, bacterial, and inflammatory diseases in febrile children remains a major clinical challenge that contributes to diagnostic uncertainty, inappropriate antimicrobial use, and suboptimal clinical management. Host blood transcriptomics offer a promising strategy to improve diagnostic precision. The present study represents the largest integrative multi-cohort pediatric study of transcriptomic biomarker discovery, validation, and confirmation reported to date, integrating harmonized public transcriptomic datasets with an independent confirmation cohort comprising well-phenotyped patients to identify parsimonious host-response signatures for differentiating viral, bacterial, and inflammatory diseases. Transcriptomic signatures were derived from an integrated retrospective microarray multi-cohort (n=1,683), independently validated in a retrospective RNA-seq cohort (n=767), and confirmed by digital PCR in an independent cohort (n=29), demonstrating reproducibility across patient populations, transcriptomic technologies, and analytical platforms. The analysis identified binary signatures and a unified multiclass classifier that consistently achieved high diagnostic accuracy across all three study phases and outperformed more than 30 published host transcriptomic signatures. Decision curve analysis showed substantially greater clinical net benefit than C-reactive protein across clinically relevant decision thresholds. These findings provide a strong foundation for clinically deployable molecular diagnostics to improve patient triage, antimicrobial stewardship, and precision medicine in childhood infections.
Lee, T. S. E.; Nguyen, L.; Forde, B. M.; Maidment, T.; Ye, S.; Henderson, A.; Playford, E. G.; Runnegar, N.; Henderson, B.; Watson, C.; Lindsay, M.; Bursle, E.; Douglas, J.; Hume, J.; Paterson, D. L.; Kidd, T.; Graves, B.; Hume, A.; Hall, M. B.; Schembri, M. A.; Beatson, S. A.; Harris, P. N. A.; Roberts, L. W.
Show abstract
OXA-48-like carbapenemases have been historically rare, however steady increases both locally and globally have warranted further investigation into their spread. Here we present the largest genomic analysis of blaOXA-181-producing bacteria in Australia to date, focusing on a single jurisdiction over seven years (2017 -- 2024). The initial investigation was prompted by an outbreak in 2017, where enhanced genomic surveillance in a single hospital identified 85 outbreak isolates related to an imported Escherichia coli ST38, carrying blaOXA-181 on an IncX3/colKP3 plasmid (previously reported as pOXA181). After four months of intensive infection control, the initial outbreak strain was eliminated. To confirm the outbreak plasmid was also contained, we collected all blaOXA-181-positive isolates from the same jurisdiction over subsequent years and sequenced with both Illumina and Oxford Nanopore Technologies to investigate clonal and mobile genetic element mediated spread. While continued surveillance post-2017 did not identify the same E. coli strain following the outbreak, pOXA181 plasmids were identified in >70% of surveillance isolates, with minimal genetic changes, which initially suggested local plasmid-mediated spread. Additional comparison to a global collection of pOXA181 plasmids found that epidemiologically unrelated pOXA181 plasmids were near identical, with no rearrangements and low, or no, single nucleotide polymorphisms. This suggests the mutation rate of pOXA-181 is incompatible with recent genomic transmission inference. This study highlights the current genomic epidemiology and drivers of blaOXA-181 and further demonstrates the necessity for detailed understanding of plasmid evolutionary rates to inform genomic surveillance.
Pagnuco, I.; Eyre, S.; Rattray, M.; Morris, A. P.
Show abstract
Type 2 diabetes (T2D) is a complex metabolic disorder characterized by hyperglycemia and insulin resistance. Although genome-wide association studies (GWAS) have identified >600 T2D risk loci, the causal genes and the relevant tissues mediating these associations remain largely unresolved. To address this challenge, we performed tissue-specific, ancestry-aware transcriptome-wide association studies (TWAS) across six T2D-relevant tissues: subcutaneous adipose, visceral adipose, brain hypothalamus, liver, skeletal muscle, and pancreas. We conducted ancestry-specific multi-tissue TWAS in European ancestry (EUR) data using summary statistics from the largest EUR GWAS (242,283 cases and 1,569,734 controls) and pre-trained gene expression prediction models derived from 689 EUR individuals from the Genotype-Tissue Expression (GTEx) Project. Conditional analyses were performed to identify independent TWAS signals. We identified 684-750 significant gene-T2D associations per tissue (P < 1.919 x 10-6), implicating both established and novel candidate genes. Among these, JAZF1 and IDE showed consistent association signals across all six tissues, whereas TCF7L2 and WSF1 exhibited heterogeneous effects restricted to a subset of T2D-relevant tissues. Conditional analyses further refined these signals to 289-322 independent TWAS signals per tissue. Together, these finding highlight substantial regulatory heterogeneity in the genetic architecture of T2D and underscore the importance of tissue context in interpreting disease-associated loci. Cross-ancestry replication of EUR-derived TWAS signals was evaluated in African American (AFA) individuals. We conducted an AFA-TWAS using summary statistics from the largest AFA GWAS (50,251 cases and 103,909 controls) in combination with gene expression prediction models trained in 111 AFA individuals from GTEx. We observed significant enrichment of EUR-derived T2D TWAS signals in the AFA TWAS across subcutaneous adipose, visceral adipose, skeletal muscle, and pancreas, whilst enrichment was weaker in liver, likely reflecting limited sample size. Overall, our findings demonstrate that integrating tissue-specific and ancestry-aware TWAS refines the identification of causal genes for T2D, with cross-ancestry replication supporting the robustness of these signals and cross-tissue analyses revealing context-specific effects. However, they also highlight the limited availability of non-EUR datasets and the need for larger, more diverse ancestry-specific transcriptomic resources.
SULAIMAN, M. A.; Oyeyemi, B. F.
Show abstract
Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.
Zhou, X.; Le, Z.; Song, P.; Xu, Q.; Chen, M.; Liu, X.; Cao, M.; Zhan, S.; Liu, Y.; Zhang, L.
Show abstract
Background: Inflammation and the tumor immune microenvironment contribute to lung adenocarcinoma (LUAD) progression, but the relationship among inflammation-linked transcriptional heterogeneity, patient survival, and immune-state variation remains incompletely defined. Objective: We aimed to identify inflammation-associated LUAD subtypes, derive a parsimonious survival-stratification signature, and characterize its immune and pathway context across public transcriptomic cohorts. Methods: Expression profiles and clinical data were obtained from TCGA-LUAD, GTEx normal lung, and GEO datasets GSE11969, GSE30219, GSE31210, and GSE40791. A curated set of 596 inflammation-related genes was used for consensus clustering. Differential-expression analysis, functional enrichment, univariate Cox regression, and LASSO-Cox modeling were integrated to construct a gene-expression risk score. The prognostic dataset comprised 730 cases and was randomly divided into training (n=502) and internal-validation (n=228) sets; 85 GSE30219 cases formed an external-validation cohort. Immune-cell enrichment, gene set enrichment analysis (GSEA), gene set variation analysis (GSVA), and pan-cancer analyses were used for biological contextualization. Results: The LUAD-versus-control comparison identified 1,305 differentially expressed genes, including 498 upregulated and 807 downregulated genes. Consensus clustering resolved two inflammation-associated subtypes and 67 subtype-associated genes, of which 64 were higher and 3 were lower in Cluster 1 relative to Cluster 2. Thirty-three genes overlapped between the tumor-control and subtype contrasts. LASSO-Cox regression selected CHRDL1, FDCSP, CXCL13, CYP4B1, and S100P. The 1-, 3-, and 5-year areas under the time-dependent receiver operating characteristic curve were 0.6625, 0.6581, and 0.6658 in the training set; 0.7422, 0.6537, and 0.6761 in internal validation; and 0.6560, 0.6387, and 0.6753 in external validation. Risk groups differed across multiple T-cell, B-cell, natural-killer-cell, myeloid, dendritic-cell, macrophage, and granulocyte signatures. Positive GSEA signals included cell cycle (normalized enrichment score [NES]=2.67; adjusted P=1.42 x 10-), DNA replication (NES=2.52; adjusted P=2.52 x 10-), and mismatch repair (NES=2.20; adjusted P=1.77 x 10-). Conclusions: The five-gene expression score separated LUAD survival groups and captured coordinated proliferative and immune transcriptional states. Its moderate discrimination supports further biological and clinical validation rather than immediate clinical application.
Hodel, F.; Thorball, C. W.; Haefliger, D.; Cerutti, L.; Cattaneo, P.; Howald, C.; Männik, K.; de La Harpe, R.; Samer, C. F.; Xenarios, I.; Fellay, J.; Girardin, F. R.
Show abstract
Background. Pharmacogenetic (PGx) testing can guide drug prescribing but remains limited by the genomic assay used. Genotyping arrays are widely implemented yet limited to predefined variants, whereas low-pass whole-genome sequencing (LP-WGS) is not constrained by fixed probe design and may provide broader PGx variant availability after imputation. Methods. We compared Illumina Global Screening Array (GSA) v3 with ~1x LP-WGS for PGx profiling in 500 hospital biobank participants with electronic health record evidence of exposure to pharmacogenetically actionable drugs and reported adverse drug reactions. Concordance was evaluated genome-wide, at 20 actionable pharmacogenes for PharmCAT-derived star alleles and metabolizer phenotypes, and for HLA alleles. Results. Genome-wide concordance between imputed array and LP-WGS data was high (median 99.63%; interquartile range, 99.59%-99.64%). For pharmacogenetically relevant variants, LP-WGS captured a larger fraction, particularly rare alleles absent from the array data, whilst maintaining high concordance at shared sites. Predicted phenotype concordance exceeded 98% for most genes, although gene-specific differences in phenotype classification were observed. LP-WGS reduced missing phenotype assignments for selected loci, particularly CYP2C19 and NAT2, by improving resolution of star-allele structure. However, in structurally complex or incompletely characterized genes such as CYP2C9 and CYP2D6, broader variant recovery increased indeterminate classifications rather than consistently improving clinical interpretability. For HLA loci, concordance varied by imputation strategy, with SNP2HLA performing marginally better utilizing the GSA array compared to the LP-WGS approach. Conclusions. Overall, LP-WGS provides broader variant coverage and improved resolution for selected pharmacogenes but did not resolve all clinically important loci. These findings support further evaluation of LP-WGS as a scalable PGx screening approach, especially where long-term genomic data reuse is a priority.
Barbosa Araujo, P. V.; da Silva Fiuza, T.; Ferraz, R. S.; Kroll, J. E.; Andrade, R. L.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; de Souza, S. J.
Show abstract
Polygenic risk scores (PRS) have emerged as a powerful tool for quantifying genetic susceptibility to complex traits and diseases. However, their calculation and interpretation require standardized data curation, robust statistical methods, and clear reporting strategies. In this work, we present an integrated pipeline designed to address these challenges. The pipeline begins with the construction of a curated genotype/phenotype database derived from public repositories, ensuring that only phenotypes with appropriate metadata, statistical distributions, and ethical suitability are retained. The final dataset comprises 2,346 phenotypes covering 38,256,468 unique SNPs. These phenotypes serve as the final analytical units for PRS calculation, risk stratification, and individual-level interpretation. The generated reports integrate sample-level results, phenotype categorization, risk classification, study references, and variant tables, providing a structured and interpretable output for end users. Together, the curated database and reporting framework establish a comprehensive toolbox for PRS analysis, enhancing reproducibility, transparency, and usability in both research and clinical contexts.
Atteih, S. E.; Raraigh, K. S.; Wu, M.; Collaco, J. M.; Blackman, S. M.
Show abstract
Diabetes is a highly prevalent complication of cystic fibrosis (CF), affecting 50% of adults with CF and over 80% of those with exocrine pancreatic insufficiency (PI) by age 50 years. Development of cystic fibrosis-related diabetes (CFRD) is associated with increased morbidity and mortality mostly due to advancement of chronic obstructive lung disease. Highly effective modulator therapy (HEMT), using precision medications targeting the cystic fibrosis transmembrane conductance regulator (CFTR), improves CFTR function and CF lung disease, but its impact on diabetes pathogenesis remains uncertain. We sought to determine whether two types of HEMT, ivacaftor and elexacaftor/tezacaftor/ivacaftor (ETI), alter diabetes prevalence in two large cohorts of individuals with CF and exocrine PI. For comparison, a non-highly-effective modulator, lumacaftor/ivacaftor (LUM/IVA), was also assessed. Data were provided by the CFTR2 project, a multinational CF registry (for ivacaftor and LUM-IVA), and by the CF Genome Project (CFGP), a predominantly US-based CF cohort (for ETI). Among 32,753 individuals with CF (2,803 treated), ivacaftor was associated with reduced diabetes prevalence (age-adjusted OR=0.55). In contrast, lumacaftor/ivacaftor (not highly effective) was not associated with diabetes prevalence (n=32,749). Among 2,854 individuals with CF (2,458 treated), ETI was associated with reduced diabetes prevalence (age-adjusted OR=0.47). Overall, HEMT (ivacaftor and ETI) was associated with a 25-39% reduction in diabetes prevalence in CF, while a non-highly-effective modulator (lumacaftor/ivacaftor) showed no difference. Precision targeted amelioration of CFTR dysfunction can delay onset of diabetes in a high-risk CF population.
Wan, X.; Ji, L.; Han, S.; Lin, Z.; Zeng, Y.; Ming, M.; Gan, W.; Duan, X.; Lu, H.; Shen, J.
Show abstract
Bacterial defense systems against bacteriophages are critical for bacterial genome stability and fitness, yet their distribution and epidemiological relevance in Klebsiella pneumoniae remain unexplored. We aimed to characterize defense system signatures across major K. pneumoniae clonal lineages and evaluate their utility as genomic signatures for tracking clonal dissemination. Here, we analyzed 6,346 genomes to characterize population structure, defense system signatures, and their coevolution with resistance determinants and plasmid backbones. PopPUNK clustering resolved 146 lineages that stratified into major clonal lineages (G-ST-KL combinations) with distinct resistance and virulence profiles. DefenseFinder and PADLOC identified 320 distinct defense systems, revealing that each clonal lineage harbors a unique defensotype characterized by systematic module replacement rather than stochastic gene loss. Co-occurrence networks further showed that these systems are organized into lineage-specific functional modules, with contrasting architectures even among lineages sharing the same sequence type. Integration of defense, resistance, and plasmid data uncovered strong lineage-specific associations, whereby broad-host-range plasmid backbones acquired distinct defense-resistance payloads in different clonal backgrounds. Geographic and host-niche analyses demonstrated that defense system distribution reflects clonal lineage expansion rather than independent geographic selection, and analysis of 689 Chinese genomes confirmed vertical inheritance of lineage-specific signatures along transmission chains. Collectively, defense systems in K. pneumoniae are organized into lineage-specific defensotypes shaped by synergistic modules and clonal evolutionary dynamics. Defense system profiling provides an additional layer of epidemiological resolution beyond conventional typing and offers practical utility for genomic surveillance, particularly in resource-limited settings where PCR-based detection of conserved systems could serve as a rapid proxy for identifying high-risk K. pneumoniae clones.
van Oosten, D.; Beele, P.; Wang, B.-n.; Plasmans, S. J.; Wolthuis, N.; van den Berg, K.; Blom, M. P. T.; Meyjes, M.; van der Schoot, N. D.; Vergunst-Bosch, H.; Kok, A. R.; van der Ven, L. J.; van Es, M. A.; van den Berg, L. H.; Veldink, J. H.; van Rheenen, W.
Show abstract
Importance: With emerging gene-targeted therapies in amyotrophic lateral sclerosis (ALS), gene discoveries and genetic diagnoses provide a crucial path to treatment. Pathogenic variants with moderate effect or incomplete penetrance, however, remain unidentified in genome-wide association studies and can appear sporadic in small modern-day pedigrees. Lack of recognition of familial clustering of ALS, in turn, limits opportunities for gene discovery, genetic diagnosis, risk counseling, and treatment. Objective: To determine the power of automated reconstruction of extended pedigrees, integrating archive records and genetic relatedness, in gene-discovery studies. Design: Retrospective observational study of Dutch ALS patients with the C9orf72 hexanucleotide repeat expansion (HRE), combining clinical family history, civil records, and genome-wide genotyping for relatedness and identity-by-descent (IBD) inference. Setting: National, population-based ALS cohort from the Netherlands and digitized population archives enabling systematic reconstruction of extended pedigrees. Participants: Individuals with ALS and a confirmed C9orf72 HRE. Participants must have provided a clinical family history and traceable Dutch ancestry documented in population archives. Main Outcomes and Measures: The primary outcome was the proportion of C9orf72 HRE carriers with newly identified (distant) relatives with ALS compared with clinical family history. The secondary outcome was the precision of IBD-based methods to fine-map the C9orf72 HRE. Other outcomes included phenotypic similarities between distantly related patients. Results: Among 238 C9orf72 HRE carriers, 91 could be included in one of 39 extended pedigrees dating back to ~1800, with relationships up to the eighth degree of relatedness. Compared with clinical family history alone, our approach increased the number of identified relationships by 2.5-fold. Genome-wide IBD analysis revealed shared haplotypes encompassing the C9orf72 HRE in 94% of pedigrees by [≥]7 meioses in 25.7-127.8 centimorgans total IBD shared. Conclusions and Relevance: Large-scale interrogation of archives facilitates reconstruction of extended pedigrees for ALS patients carrying the C9orf72 HRE. This combined genealogical-genetic approach supports the reclassification of apparently sporadic cases, facilitates the discovery of new disease-causing variants in ALS, and is generalizable to other late-onset neurodegenerative diseases. Automated pedigree reconstruction from genealogical data and visualization in an interactive databrowser are implemented in the open-source Mangrove software.
Shih, K. Y.; Brandman, O.; Winslow, M. M.; Petrov, D. A.
Show abstract
Tumor mutational burden (TMB) shapes tumor transcriptional state, but studies typically describe this response as an average effect pooled across cancer types. Whether that average reflects a consistent response present within individual cancer types, or is an artifact of merging heterogeneous, tissue-specific responses, remains unresolved. Here we analyze ~9,100 tumors across 32 TCGA cancer types to test whether the transcriptional response to TMB is genuinely consistent across tissues. We construct a TMB axis score from TMB-associated genes upregulated with increasing TMB, yielding a sample-level measure of response strength, and subsequently decompose it at the component and pathway/complex levels. The pooled transcriptional response to TMB stays largely consistent within each cancer type, and no single cancer is driving the pooled signal. This consistency was also observed at the component and pathway/complex levels. These findings support TMB as a promising tissue-agnostic signature, with implications for tissue-agnostic therapeutic targeting.
Desboeufs, N.; Leary, P.; Zhao, C.; Kollar, S.; Chan, L. K.; Planas-Paz, L.; Fitsche, A.; Schmidt, A.; Prutek, F.; Baumann, K. R.; Schneebeli, S.; Dettwiler, S.; Dona, F.; Akpinar, R.; Terracciano, L. M.; Piscuoglio, S.; Di Tommaso, L.; Wild, K.; Summermatter, L.; Kobe, A.; Puippe, G. D.; Leblond, A.-L.; Endhardt, K.; Ng, C. K. Y.; Nuciforo, S.; Heim, M. H.; Fritsch, R.; Pauli, C.; Kremer, A. E.; Lopes, M.; Weber, A.
Show abstract
Background: To date, no precision oncology approach has been established for HCC. Despite the diverse underlying causes, HCC development exhibits a uniform pathophysiology characterised by chronic hyper-proliferation, resulting from hepatocyte apoptosis and compensatory liver regeneration. This chronic hyper-proliferative pressure, termed regeneration stress, drives genomic instability during HCC onset, yet its therapeutic potential remains poorly explored. This study aimed to identify targetable vulnerabilities tied to regeneration stress and establish clinically applicable markers for treatment stratification. Methods: Weighted gene co-expression network analysis (WGCNA) was applied on external bulk RNA-seq datasets to define a LIVer REgeneration Stress Signature (LIVRESS). The signature was functionally validated using HCC patient-derived organoids (HCC-Org), and vulnerabilities were mapped using mid-throughput drug screening, single-molecule and single-cell assays, and multi-omic integration. Results: High LIVRESS scores, characterised by enrichment in replication, mitotic and DNA damage repair pathways, identified a subset of HCC patients with aggressive disease and poorer survival across aetiologies. HCC-Org with high LIVRESS scores displayed exquisite sensitivity to multiple inhibitors of the checkpoint kinase ATR. Although HCC-Org models exhibited a baseline reduction in replication fork speed, sensitivity to ATR inhibitor (ATRi) was decoupled from replication fork dynamics and rather linked to intrinsic mitotic instability. ATR inhibition triggers mitotic failure and apoptosis in LIVRESSHigh HCC-Org. This killing effect was significantly potentiated by combining ATRi with PARPi or WEE1i. Multi-omic integration identified KPNA2 as a surrogate biomarker of ATRi sensitivity. Conclusion: Our findings demonstrate that a subset of HCC-Org, characterised by high liver regeneration-associated stress, is vulnerable to ATRi-based therapies. By focusing on a comprehensive regenerative stress model, we establish a framework to stratify HCC patients and implement biomarker-driven, ATR-based therapies for HCC patients with advanced disease. Impact and implications: Regeneration stress is a key factor that drives genomic instability in HCC, providing a basis for the LIVRESS to identify patients dependent on ATR-mediated checkpoints. These findings reveal a conceptual shift for researchers and trialists: ATRi efficacy is decoupled from replication fork dynamics and instead leverages mitotic fragility. Practically, the LIVRESS and its IHC surrogate marker (KPNA2) offer a scalable roadmap for physicians to improve patient stratification in ATRi-based precision oncology trials. While requiring prospective validation, these results pave the way toward biomarker-driven therapies for advanced HCC.
Elrod, J. K.; Sanyal, A.; Hutchins, T.; Townes, F. W.; Torok, K. S.
Show abstract
Juvenile systemic sclerosis (jSSc) is a rare autoimmune disease marked by skin fibrosis and multi-organ involvement. Autologous stem cell transplantation (ASCT) is an emerging therapy for severe, treatment-refractory jSSc, but its effects on immune cell dynamics remain poorly understood. PBMCs were collected from three patients with jSSc before ASCT and at 6, 12, and 24 months post-ASCT. Patient and healthy control samples were profiled using cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq). We focused on monocytes, given their role in fibrosis-promoting inflammation. To detect longitudinal trends, pseudobulked gene expression (log scale) was regressed against time since ASCT. This approach identified widespread changes in jSSc monocytes, including decreased expression of systemic sclerosis-linked genes, such as SERPINE1. On the pathway level, NF-{kappa}B-associated inflammatory signaling was elevated in jSSc monocytes at baseline relative to healthy controls and decreased progressively post-ASCT. Genes related to mitochondrial function and oxidative phosphorylation progressively increased in expression after ASCT, suggesting a shift in metabolic state. Compositional changes in monocyte subpopulations were also identified and may have contributed to longitudinal gene expression patterns. Together, these findings characterize the dynamic immune changes in jSSc following ASCT and highlight a widely applicable longitudinal modeling framework for single-cell data.
Maffa, S.; Boyle, I. A.; Ward, L.; Colgan, W. N.; Borck, P.; Simerzin, A.; Adeagbo, A.; Olajide, O.; Wie, S.; Liang, H.; Wienand, K.; Shibue, T.; Ray, J.; Paolella, B.; Campbell, C. D.; Vazquez, F.; Dempster, J. M.
Show abstract
Background CRISPR-mediated viability assays in diverse cancer cell lines have informed cancer biology and precision medicine, but cell fitness is not the only cancer-relevant phenotype. Gene expression profiling provides insight into cellular stress, inflammation, and differential state, while still identifying activation of cell-death pathways. Perturb-seq allows scalable functional genomics screening of expression phenotypes at single-cell resolution, however existing datasets cover only a small number of work-horse cell lines. Results We produced a proof-of-concept Perturb-seq dataset targeting 100 genes in 16 diverse cancer cell lines. In the process, we established methods to address single-cell technical artifacts, identified Cas9-mediated chromosomal aberrations and assessed screen quality. Even with a limited library, we observed common signatures of deleting essential genes as well as context-specific responses based on intrinsic genomic properties of the models. For example, we inferred a previously undescribed relationship between dependence on the ER-golgi transport gene immediate early response 3 interacting protein 1 (IER3IP1) and oxidative stress, demonstrating the potential of integrated Perturb-seq for hypothesis generation. Conclusions We established a framework for building a comprehensive map of post-perturbational transcriptional phenotypes using parallel Perturb-seq experiments across multiple cell lines. We demonstrated that integrated Perturb-seq experiments spanning diverse contexts enable hypotheses about gene function specific to tissue types or cancer subtypes - suggesting large-scale, genome-wide datasets would offer invaluable insight into the highly context-dependent nature of cancer biology.
NG, I. C.-F.; WONG, I. T.-F.; LEUNG, J. S.-L.; LEE, L.-K.; LAM, A. Y.-T.; TONG, H.-C.; CHAN, S.-K.; Wong, C.-Y.; LEE, A. W.-T.; TAM, W.-Y.; ZHANG, J.-Y.; HILL, E. M.; HUNG, M.-F.; YAU, M. C.-Y.; WONG, R. C.-W.; CHENG, J. C.-K.; TSE, C. W.-S.; LAM, J. Y.-W.; CHOW, V. C. Y.; CHAU, S. K.-Y.; Chow, F. W.-N.; LEUNG, P. H.-M.; Siu, G. K. H.
Show abstract
Carbapenem-resistant Escherichia coli (CR-E. coli) is an emerging One Health threat, but recent shifts in predominant lineages and genomic links between clinical and food reservoirs in Hong Kong remain poorly defined. We analyzed 271 CR-E. coli isolates from four hospitals (2022-2026) and 585 isolates recovered from 4,917 retail food samples (2022-2025). Isolates underwent antimicrobial susceptibility testing, whole-genome sequencing, multilocus sequence typing, resistance-gene and plasmid profiling, core-genome SNP phylogenetics, and comparative genomics. Food isolates were mainly from raw pork (268/585, 45.8%) and raw chicken (231/585, 39.5%). blaNDM-5 was detected in 527/585 (90.1%) food and 241/271 (88.9%) clinical isolates. ST69 was the most frequent defined sequence type in both collections, representing 44/585 (7.5%) food and 36/271 (13.3%) clinical isolates, in contrast to the heterogeneous lineages and carbapenemases previously reported in Hong Kong. Applying a predefined [≤]50-pairwise-SNP threshold for close genomic relatedness, core-genome phylogeny of 80 ST69 isolates identified two major mixed-source clusters collectively comprising 28 clinical and 27 food isolates. Clustered isolates showed similar antimicrobial resistance profiles, carried blaNDM-5 and blaTEM-1, and were associated with IncI1 MLST | ST136 plasmids. Comparative analyses showed >99.85% average nucleotide identity and broad conservation of the blaNDM-5-associated plasmid backbone across sources. These findings indicate the emergence of blaNDM-5-carrying ST69 as a prominent CR-E. coli lineage in Hong Kong and demonstrate close genomic relatedness between selected clinical and retail food isolates. Although transmission direction have not been inferred yet, the findings support integrated One Health surveillance and source-tracing across clinical, food, animal, and environmental sectors.
Altman, G. N.; Jadhav, B.; Garg, P.; Shadrina, M.; Manigbas, C. A.; Lee, W.; Kandoi, S.; Martin-Trujillo, A.; Sharp, A. J.
Show abstract
Tandem repeat expansions (TREs) cause over 50 neurological conditions, yet their contribution to neurodegenerative disease risk at a population scale remains incompletely characterized. We performed a TRE association study across 6,539 short tandem repeat loci in 276,411 individuals from the UK Biobank and 44,370 individuals from the All of Us Research Program, using two composite neurodegenerative phenotypes to increase statistical power and capture pleiotropic effects. Meta-analysis across the two cohorts identified associations at eight established pathogenic TRE loci, including C9orf72, DMPK, HTT, ATXN2, ATXN3, CACNA1A, CNBP, and PPP2R2B, recovering known disease-associated expansions from short-read sequencing data at biobank scale. We also identified candidate associations at three additional loci. An intronic AATAA expansion in DAPK1 reached significance (q = 0.0045), with fine-mapping and conditional analysis supporting the repeat as the likely variant underlying the association. An intronic ATTTT expansion in ANK3 (q = 0.034) was observed exclusively in individuals of African and Latino/admixed American ancestry, underscoring the importance of ancestrally diverse cohorts for genetic discovery. An exonic polyalanine expansion in RPL14 was also significant (q = 0.039), where longer alleles were consistently associated with reduced RPL14 expression across independent datasets. Together, these findings identify candidate risk loci for neurodegenerative disease that may expand the contribution of TREs to neurodegenerative disease beyond known repeat expansion disorders.